Papers with end-to-end training

31 papers
DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) with web search capabilities show significant potential for deep research.
Approach: They introduce a framework for end-to-end training of LLM-based deep research agents . they implement a specialized multi-agent architecture where browsing agents extract relevant information from various webpage structures.
Outcome: The proposed framework improves on open-domain research tasks by 28.9 points over prompt engineering and 7.2 points over RAG-based RL agents.
Autoregressive Entity Generation for End-to-End Task-Oriented Dialog (2022.coling-1)

Copied to clipboard

Challenge: Task-oriented dialog systems require external knowledge base to generate a response . current systems require scanning the KB at each turn, which is inefficient when the kb scales up .
Approach: They propose to generate entity autoregressively before leveraging it to guide response generation.
Outcome: Experiments on MultiWOZ 2.1 single and CAMREST show that the proposed system generates more high-quality and entity-consistent responses in an end-to-end manner.
Sample, Translate, Recombine: Leveraging Audio Alignments for Data Augmentation in End-to-end Speech Translation (2022.acl-short)

Copied to clipboard

Challenge: End-to-end speech translation relies on data that pair source-language speech inputs with corresponding translations.
Approach: They propose a method that augments transcriptions by sampling from suffix memory and translating them into target languages.
Outcome: The proposed method delivers up to 0.9 and 1.1 BLEU points on top of augmentation with knowledge distillation on languages on CoVoST 2 and Europarl-ST.
Augmenting Neural Networks with First-order Logic (P19-1)

Copied to clipboard

Challenge: Existing paradigms for training neural networks require large datasets, a paper argues . we present a framework for introducing declarative knowledge to neural networks .
Approach: They propose a framework for introducing declarative knowledge to neural networks . they compile logical statements into graphs that augment a network without extra learnable parameters or manual redesign.
Outcome: The proposed framework improves on three tasks, especially in low-data regimes.
Scene Graph Parsing as Dependency Parsing (N18-1)

Copied to clipboard

Challenge: Recent studies have focused on parsing structured knowledge graphs from textual descriptions.
Approach: They propose an alternative but equivalent scene graph representation that connects to dependency parses.
Outcome: The proposed model outperforms best approaches on image retrieval applications.
Hierarchical Text Classification with Reinforced Label Assignment (D19-1)

Copied to clipboard

Challenge: Existing hierarchical text classification methods make local decisions regarding labels or ignore hierarchy information during inference.
Approach: They propose to learn a Label Assignment Policy via deep reinforcement learning to determine where to place an object and when to stop the assignment process.
Outcome: The proposed method outperforms state-of-the-art methods on five datasets and four base models and achieves an average improvement of 33.4% over flat classifiers.
Beyond Fine-tuning: Few-Sample Sentence Embedding Transfer (2020.aacl-main)

Copied to clipboard

Challenge: Fine-tuning (FT) pre-trained sentence embedding models on small datasets has been shown to have limitations.
Approach: They propose to combine embeddings from a pre-trained model with a simple sentence embeddable model.
Outcome: The proposed approach outperforms FT on small datasets with negligible computational overhead.
Self-Training with Weak Supervision (2021.naacl-main)

Copied to clipboard

Challenge: State-of-the-art deep neural networks require large amounts of labeled training data that is expensive to obtain or not available for many tasks.
Approach: They propose a weak supervision framework that leverages all available data for a given task . they leverage task-specific unlabeled data through self-training with a model that predicts pseudo-labels for instances that may not be covered by weak rules .
Outcome: The proposed framework improves on state-of-the-art datasets on six benchmark tasks.
OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model (2025.emnlp-demos)

Copied to clipboard

Challenge: Empathetic speech models are increasingly closed off, leaving details about the architecture, data and development opaque to researchers.
Approach: They propose an open-source empathetic speech-to-text model with a streaming interleaved decoding architecture and a data pipeline to enable end-to end training.
Outcome: The proposed model is open-source and transparent, with no data or data required to build it.
Self-Critique Guided Iterative Reasoning for Multi-hop Question Answering (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable reasoning capabilities, but they still face challenges in knowledge-intensive multi-hop reasoning.
Approach: They propose a method that uses self-critique feedback to guide iterative reasoning by enabling iteration and self-evaluation of its intermediate reasoning steps.
Outcome: The proposed method surpasses the previous SOTA by 8.6% on three multi-hop reasoning datasets.
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing research treats MLLMs as unified systems optimized through end-to-end training, but the impact of vision encoder’s prior knowledge is seldom investigated.
Approach: They propose a metric to quantify the effect of prior knowledge on MLLM performance by integrating prior knowledge at the vision encoder level into a training framework.
Outcome: The proposed training framework incorporates prior knowledge at the vision encoder level, and significantly boosts visual understanding capabilities of MLLMs.
Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison (2025.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been successful in NLP tasks, but there is growing interest in extending their capabilities to speech.
Approach: They propose to use dense feature prepending (DFP) to integrate speech into LLMs to enable end-to-end training with a speech encoder.
Outcome: The proposed approach does not show a clear advantage over cross-attention.
Understanding the Mechanics of SPIGOT: Surrogate Gradients for Latent Structure Learning (2020.emnlp-main)

Copied to clipboard

Challenge: Latent structure models can mitigate the error propagation and annotation bottleneck in pipeline systems, while uncovering linguistic insights about the data.
Approach: They propose a latent structure model with a pullback of the downstream learning objective.
Outcome: The proposed model outperforms the known and proposed model in the same family and yields new insights for practitioners and revealing intriguing failure cases.
UCGRec: User-Centric Graph Learning for LLM-based Sequential Recommendation (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for sequential recommendation rely primarily on item descriptions or utilize user preferences independently.
Approach: They propose a method that integrates diverse user-relevant preference signals into a unified user-centric graph and injects the graph-based knowledge into the LLM through end-to-end training with graph neural networks.
Outcome: The proposed method outperforms conventional and state-of-the-art methods on four widely used sequential real-world recommendation datasets.
Hierarchical Sketch Induction for Paraphrase Generation (2022.acl-long)

Copied to clipboard

Challenge: Existing models of paraphrase generation are based on a syntactic sketch, but prior work has included inductive bias.
Approach: They propose a method for learning decompositions of dense encodings as a sequence of discrete latent variables that make iterative refinements of increasing granularity.
Outcome: The proposed model improves on human paraphrase generation by predicting syntactic sketches at test time.
PRAM: An End-to-end Prototype-based Representation Alignment Model for Zero-resource Cross-lingual Named Entity Recognition (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to address the named entity recognition problem are limited and lack explicit optimization specific to the task.
Approach: They propose a prototype-based representation alignment model for a cross-lingual named entity recognition task using labeled source language data.
Outcome: The proposed model outperforms existing state-of-the-art methods in some challenging scenarios.
Modeling Label Correlations for Ultra-Fine Entity Typing with Neural Pairwise Conditional Random Field (2022.emnlp-main)

Copied to clipboard

Challenge: Entity typing assigns semantic types to entities mentioned in text.
Approach: They propose to use an undirected graphical model to formulate the UFET problem by combining unary potentials with a pairwise conditional random field model.
Outcome: The proposed model outperforms the existing model with little cost and is thousands of times faster than the existing neural network module.
Phrase Grounding by Soft-Label Chain Conditional Random Field (D19-1)

Copied to clipboard

Challenge: Existing methods to ground entities depend on inference or non-differentiable losses.
Approach: They propose a phrase grounding task that grounds entities to corresponding regions in an image . they use neural chain Conditional Random Fields to model dependencies among regions .
Outcome: The proposed method is based on a dataset of the Flickr30k Entities dataset.
End-to-End Training of Neural Retrievers for Open-Domain Question Answering (2021.acl-long)

Copied to clipboard

Challenge: Recent work on training neural retrievers for open-domain question answering (OpenQA) has employed both supervised and unsupervised methods.
Approach: They propose an approach of unsupervised pre-training with the Inverse Cloze Task and masked salient spans followed by supervised finetuning using question-context pairs.
Outcome: The proposed approach outperforms models like REALM and RAG in retrieval accuracy and answer extraction.
Document Hashing with Mixture-Prior Generative Models (D19-1)

Copied to clipboard

Challenge: Existing generative hashing methods only consider the use of simple priors, which limits them to further improve their performance.
Approach: They propose to use Gaussian and Bernoulli priors to generate hashing codes . they propose to cast a Gausssian latent representation into binary code .
Outcome: The proposed models outperform existing methods on a benchmark dataset using Gaussian and Bernoulli priors.
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models and diffusion models have opened new possibilities for AI-generated content . personalized cover image generation remains underexplored despite its critical role in boosting user engagement on digital platforms.
Approach: They propose a framework that integrates MLLM-based prompting with personalized preference alignment to generate high-quality, contextually relevant covers.
Outcome: The proposed framework improves image quality, semantic fidelity, and personalization, leading to stronger user appeal and offline recommendation accuracy in downstream tasks.
D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents (2026.findings-acl)

Copied to clipboard

Challenge: Graphical User Interface (GUI) agents aim to automate a wide spectrum of human tasks by emulating user interaction.
Approach: They propose a deliberative framework that leverages a fine-grained tip retrieval mechanism to inform its decision-making process.
Outcome: The proposed framework achieves SOTA among open-source general models on AndroidWorld and ScreenSpot-V2 . it leverages a fine-grained, app-specific tip retrieval mechanism to inform its decision-making process .
A Differentiable Relaxation of Graph Segmentation and Alignment for AMR Parsing (2021.emnlp-main)

Copied to clipboard

Challenge: Abstract Meaning Representations (AMR) represents sentence meaning as a directed acyclic graph.
Approach: They propose to treat alignment and segmentation as latent variables and induce them as part of end-to-end training.
Outcome: The proposed model achieves significant performance gains over a 'greedy' segmentation heuristic.
On Pursuit of Designing Multi-modal Transformer for Video Grounding (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for video grounding are not end-to-end, i.e., they rely on time-consuming post-processing steps to refine predictions.
Approach: They propose an end-to-end multi-modal Transformer model that uses two encoders and a cross-modal decoder for grounding prediction.
Outcome: The proposed model is 4.9% faster than existing models and is based on a set of encodings and decoders.
From Sub-Ability Diagnosis to Human-Aligned Generation: Bridging the Gap for Text Length Control via MarkerGen (2025.acl-long)

Copied to clipboard

Challenge: Existing methods to control text length are lacking in LCTG, posing a major limitation for practical applications.
Approach: They propose a plug-and-play approach that decomposes LCTG sub-abilities with human patterns as reference and performs detailed error analysis.
Outcome: The proposed method significantly improves LCTG across various settings, exhibiting outstanding effectiveness and generalizability.
In-Image Neural Machine Translation with Segmented Pixel Sequence-to-Sequence Model (2023.findings-emnlp)

Copied to clipboard

Challenge: In-Image Machine Translation (IIMT) aims to convert images containing texts from one language to another.
Approach: They propose an end-to-end model instead of the traditional cascade methods which use optical character recognition followed by neural machine translation and text rendering.
Outcome: The proposed model outperforms both cascade methods and current model in translation quality and robustness across various dimensions.
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing backdoor models rely on visual inputs for instruction parsing, rendering the perception pathway a critical attack surface.
Approach: They propose an Adaptive Threat-Aware Adversarial Tuning framework that detects and decouples the optimal gradient decoupling strategy based on the adversary's capabilities.
Outcome: The proposed framework achieves a highly robust targeted attack success rate while maintaining extreme stealthiness with a 5% poisoning rate.
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention (2025.acl-long)

Copied to clipboard

Challenge: Long-context modeling is crucial for next-generation language models, but high computational cost of standard attention mechanisms poses significant computational challenges.
Approach: They propose a natively trained Sparse Attention mechanism that integrates algorithms with hardware-aligned optimizations to achieve efficient long-context modeling.
Outcome: The proposed model maintains or exceeds Full Attention models across general benchmarks, long-context tasks, and instruction-based reasoning.
RG-VQA: Leveraging Retriever-Generator Pipelines for Knowledge Intensive Visual Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to improve the reasoning capabilities of VQA systems are limited due to complexity of graph neural networks and end-to-end training.
Approach: They propose a method to integrate Dense Passage Retrievers with Vision Language Models to boost the reasoning capabilities of VQA systems.
Outcome: The proposed method outperforms human accuracy and GPT-4 in the ScienceQA dataset.
Towards Autonomous Tool Utilization in Language Models: A Unified, Efficient and Scalable Framework (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in tool learning for large language models have led to a new trend to allow LLMs to leverage external tools.
Approach: They propose a framework for fine-tuning language models that categorizes queries into three different types . they also introduce an "instruct, execute, and reformat" strategy specifically designed for efficient data annotation .
Outcome: The proposed framework surpasses open-source language models and GPT-3.5/4 on multiple evaluation metrics.
D-RAG: Differentiable Retrieval-Augmented Generation for Knowledge Graph Question Answering (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to Knowledge Graph Question Answering (KGQA) use Retrieval-Augmented Generation (RAG) but subgraph selection process is non-differentiable, preventing end-to-end training of the retriever and the generator.
Approach: They propose a Differentiable RAG approach that optimizes the retriever and the generator for KGQA.
Outcome: The proposed approach outperforms state-of-the-art approaches on WebQSP and CWQ.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations